With the acceleration of large model iterations, compute power and engineering have become the ultimate challenges. Li Yuxuan from Mianbi Intelligence explained the self-developed framework ForgeTrain in the 'AI4AI Fermentation Night' event, showcasing the paradigm shift of 'AI Manufacturing AI,' breaking through compute bottlenecks and engineering barriers through the underlying framework.
Due to global cloud capacity shortage, Google restricts Meta's Gemini access. The model was key for Meta's content moderation, outperforming its own Llama. Despite $20B quarterly cloud revenue, expansion lags behind surging AI inference demand.....
The efficiency of large language model inference has made a breakthrough. Tsinghua University and Moonshot AI jointly proposed a new architecture called "Prefill-as-a-Service," which splits the inference process into two stages: prefilling and decoding, and optimizes the allocation of computing resources, effectively solving hardware limitations and significantly improving model service performance.
Anthropic is evaluating the development of its own AI chip to address the surge in demand for the Claude model in 2026, enhance control over computing power, and reduce reliance on external suppliers. The company's annualized revenue has exceeded $30 billion, driving its strategic transformation with strong performance.
Google
$0.49
Input tokens/M
$2.1
Output tokens/M
1k
Context Length
Alibaba
-
Tencent
Iflytek
$2
$0.8
32
Baidu
$1
$4
64
Minimax
$1.6
$16
Openai
$21
$84
128
Chatglm
The MCP server for Dyson Sphere Program, which uses AI to assist in analyzing factory bottlenecks, power grid status, and logistics efficiency, and serves as a bridge between game data and natural language interaction.